NSF PAR Search | NSF Public Access Repository

Guardagent: Safeguard LLM agents via knowledge-enabled reasoning

Xiang, Zhen; Zheng, Linzhi; Li, Yanjie; Hong, Junyuan; Li, Qinbin; Xie, Han; Zhang, Jiawei; Xiong, Zidi; Xie, Chulin; Bastian, Nathaniel D (July 2025, ICML 2025 Workshop on Computer Use Agents)

The rapid advancement of large language model (LLM) agents has raised new concerns regarding their safety and security, which cannot be addressed by traditional textual-harm-focused LLM guardrails. We propose GuardAgent, the first guardrail agent to protect other agents by checking whether the agent actions satisfy safety guard requests. Specifically, GuardAgent first analyzes the safety guard requests to generate a task plan, and then converts this plan into guardrail code for execution. In both steps, an LLM is utilized as the reasoning component, supplemented by in-context demonstrations retrieved from a memory module storing information from previous tasks. GuardAgent can understand different safety guard requests and provide reliable code-based guardrails with high flexibility and low operational overhead. In addition, we propose two novel benchmarks: EICU-AC benchmark to assess the access control for healthcare agents and Mind2Web-SC benchmark to evaluate the safety regulations for web agents. We show that GuardAgent effectively moderates the violation actions for two types of agents on these two benchmarks with over 98% and 83% guardrail accuracies, respectively.

Full Text Available

Search for: All records